Papers by Jong Inn Park

3 papers
How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMs (2025.acl-long)

Copied to clipboard

Challenge: Large language models exhibit increasingly sophisticated linguistic capabilities, yet the extent to which these models reflect human-like cognition versus advanced pattern recognition remains an open question.
Approach: They conduct a series of targeted experiments to assess whether LLMs construct semantic representations and pragmatic inferences in a human-like manner.
Outcome: The proposed framework can be used to assess the cognitive and linguistic capabilities of large language models (LLMs).
Benchmarking Cognitive Biases in Large Language Models as Evaluators (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been shown to be effective as automatic evaluators with simple prompting and in-context learning.
Approach: They assemble 16 Large Language Models and evaluate their outputs by preference ranking . they introduce a cognitive bias benchmark to measure six different cognitive biases in LLM evaluation outputs.
Outcome: The proposed model is biased on the CoBBLer benchmark, indicating that machine preferences are misaligned with humans.
Strong Memory, Weak Control: An Empirical Study of Executive Functioning in LLMs (2026.eacl-long)

Copied to clipboard

Challenge: Working memory is a critical component of human intelligence and executive functioning . it is correlated with performance on various cognitive tasks, including fluid intelligence .
Approach: They apply working memory tasks to large language models to estimate working memory capacity . they find that LLMs exceed normative human scores, but not executive functioning benchmarks .
Outcome: The proposed models do not show higher performance on executive functioning tasks or problem solving benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations